Back

Journal of Computational Biology

SAGE Publications

Preprints posted in the last 30 days, ranked by how well they match Journal of Computational Biology's content profile, based on 48 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
Quantifying the Rearrangement Complexity of Pangenomes

Bohnenkaemper, L.; Stoye, J.

2026-08-29 bioinformatics 10.64898/2026.08.27.747493 medRxiv
Top 0.1%
38.1%
Show abstract

The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.

2
Efficient Game-Theoretic Explanations for Tree-Based Ensembles via Owen Values

Koh, H.

2026-08-20 bioinformatics 10.64898/2026.08.12.744440 medRxiv
Top 0.1%
9.5%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWShapley-value-based explanations, notably SHAP (SHapley Additive exPlanations), have gained prominence as a principled game-theoretic framework for local explanations and global feature importance. While exact Shapley value computation is exponential in feature count, TreeExplainer exploits the recursive structure of decision trees to achieve polynomial-time computation for tree-based ensembles. In many scientific applications, however, features are naturally organized into a priori groups reflecting domain knowledge, requiring explanations both across and within groups. The Owen value extends the Shapley value through a two-stage allocation rule that incorporates group structure while preserving fairness properties; yet, efficient algorithms for its computation remain limited. In this paper, we propose exact and Monte Carlo algorithms for computing Owen values in tree-based ensembles by combining hierarchy-guided group aggregation with tree-aware dynamic programming. The exact algorithm computes Owen values without sampling under the path-dependent characteristic function, which approximates the conditional expectation, whereas the Monte Carlo algorithm provides a scalable approximation that is unbiased for any prespecified sampling budget and converges almost surely as the sampling budget increases. We also provide global importance measures and visualization tools for structured, multi-resolution explanations. The proposed algorithms and tools are collectively referred to as TreeOwen. Through simulation experiments, we demonstrate the numerical accuracy and substantial computational gains of TreeOwen. We illustrate its practical utility using immunotherapy metagenomic data, showing how microbial genera (groups) and species (features) contribute to patient recovery.

3
Moirai: single-cell trajectory inference grounded in gene-level expression dynamics

Fijn, A. H. B.; S. Jeuken, G.

2026-08-11 bioinformatics 10.64898/2026.08.05.742709 medRxiv
Top 0.1%
9.5%
Show abstract

Underlying the development of multicellular organisms is the process of cell differentiation, which is governed by the concerted and sequential change in gene expression. Various methods have been developed that employ scRNA-seq data to infer the position of a cell along a pseudo-temporal axis and identify relevant genes involved in the process. These trajectory inference methods typically rely on global transcriptomic changes and mathematical methods. However, overemphasis on large-scale transcriptomic changes may impair sensitivity to identify branching points and convergent trajectories, which are rather governed by small-scale transcriptional events. Motivated by this, we developed Moirai, a graph-based trajectory inference method that identifies gene expression patterns that change dynamically over a developmental continuum and leverages these to define a common pseudotime axis between all cells. In doing so, Moirai shifts the focus to individual gene dynamics, which enhances its ability to detect putative branching points that are masked by global transcriptomic similarities. We apply Moirai to four developmental datasets, where we demonstrate its ability to recover gene expression patterns of genes with a known involvement in the respective developmental process, motivating their use for defining a cells pseudotime. We furthermore show that Moirai can robustly infer gene expression patterns across different embedding approaches, highlighting the value of moving the focus of the inference process to the small-scale transcriptional dynamics.

4
A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction

Hua, X.; Grimaud, G. M.

2026-08-18 bioinformatics 10.64898/2026.08.14.744831 medRxiv
Top 0.1%
6.5%
Show abstract

Accurate enzyme annotation remains a major bottleneck in translating rapidly growing protein sequence data into biological knowledge. Enzyme Commission (EC) prediction is particularly challenging because enzyme functions are organized hierarchically, annotations are often imbalanced across classes, and sequence similarity alone may be insufficient to resolve functional differences. To address these challenges, we developed ESM-ECForest, a two-stage framework that combines protein embeddings generated by the pretrained language model ESM-2 (Evolutionary Scale Modeling 2) with Random Forest classifiers. The first stage distinguishes enzymes from non-enzymes, whereas the second assigns one or more EC numbers to proteins predicted to be enzymatic. On an external benchmark comprising 25,778 protein sequences, ESM-ECForest achieved the highest weighted F1 score among the evaluated methods at all four EC levels, decreasing from 0.94 at Level 1 to 0.90 at Level 4. The largest relative improvements were observed for lyases (EC 4), ligases (EC 6), and translocases (EC 7), although EC 6 and EC 7 remained the most difficult classes internally. Visualization of the ESM-2 embedding space using Uniform Manifold Approximation and Projection (UMAP) revealed clustering patterns consistent with enzyme functional relationships, indicating that biologically relevant information is retained in the pretrained representations prior to supervised classification. These results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation. By combining large-scale sequence representations with a lightweight supervised classifier, ESM-ECForest provides a scalable approach for EC prediction and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.

5
FuncSeek: Multi-PLM contrastive learning for protein functional similarity search

Cloete, L. J.; Patterton, H. G.

2026-08-13 bioinformatics 10.64898/2026.08.07.738656 medRxiv
Top 0.1%
5.5%
Show abstract

Below 30% pairwise sequence identity, alignment-based methods struggle to reliably distinguish true homologs from chance (Rost 1999), and enzyme function prediction degrades accordingly: on proteins in this regime, even advanced methods (CLEAN) achieves only 55.1% accuracy at full EC specificity on the CARE benchmark (Yang et al. 2024). To this end Protein Language Models (PLMs) have gained favor as alternatives. However, PLMs often encode only a subset of the biology (Heinzinger et al. 2024, Lin et al. 2023), whereas the understanding of enzyme function requires among other things a combination of sequence, structure and functional-context simultaneously (Ribeiro et al. 2023). In this work, we describe FuncSeek, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities). Using SwissProt data, these 2816-D embeddings are labeled with Enzyme Commission numbers (EC) and are trained through a supervised contrastive head into a 256-D space. FuncSeek attains 64.6% nearest-neighbour EC4 accuracy on the CARE out-of-distribution benchmark set (ood30; proteins below 30% identity to training set), outperforming CLEAN (55.1%) and Diamond BLASTp (51.4%), and obtains 93.7% nearest-neighbour EC4 accuracy on the promiscuous, multi-functional enzymes benchmark (CLEAN, 69.4%). We also show that the learned representations transfer without retraining to the TrEMBL database, achieving 97.3% nearest-neighbour EC4 accuracy on a 8,031 BRENDA-validated enzyme set (Schomburg et al. 2004), never seen during training. Because only projected embeddings are stored in the target index, and function is inferred from an annotated reference set, we propose this paradigm for rapidly searching extremely large metagenomic databases, bypassing costly sequence alignment and annotation pipelines.

6
ProtFinder: An efficient machine learning framework for protein model selection on real data

Nguyen Huy, T.; Dong, Y.; Ly-Trong, N.; Vinh, L. S.; Minh, B. Q.

2026-08-07 evolutionary biology 10.64898/2026.08.04.742760 medRxiv
Top 0.1%
5.4%
Show abstract

Model selection is a fundamental step in phylogenetic analysis that determines the best-fit model of sequence evolution for a given multiple sequence alignment. Popular model selection methods, such as ModelFinder, rely on statistical information criteria, such as the Bayesian Information Criterion (BIC) or the Akaike Information Criterion (AIC). However, these approaches are computationally expensive and the use of information criteria has been the subject of ongoing discussion. Recently, machine learning has emerged as a promising approach for phylogenetic model selection in both nucleotide and protein sequence analyses. ModelDetector is currently the only machine learning-based method for amino acid substitution model selection. However, because ModelDetector was trained on simulated data, it does not perform well on real datasets. Another limitation is that it does not support different rate heterogeneity across sites (RHAS) models. To overcome these limitations, we introduce ProtFinder, an efficient machine learning framework for protein model selection that predicts amino acid substitution models, RHAS models, and amino acid frequency models. To enable ProtFinder to work with real datasets, we employed a transfer learning strategy consisting of three stages: (1) initial training on large-scale simulated data, (2) joint training on both simulated and real data, and (3) final fine-tuning using real data only. Experimental results show that ProtFinder outperformed ModelDetector in amino acid substitution model selection. ProtFinder achieved comparable accuracy to the maximum likelihood method ModelFinder for substitution model selection on medium and large MSAs. It performs slightly better than ModelFinder in RHAS model selection and substantially outperforms it in amino acid frequency model determination. Notably, ProtFinder is up to 1,400 times faster than ModelFinder in terms of inference time, making it particularly suitable for medium and large datasets.

7
Principal Genes: A PCA-based approach to highly variable genes selection for scRNA-Seq analysis

Kakwambi, E. D.; Nguyen, T.; Kapoor, S.; Moussa, M. R.

2026-08-13 bioinformatics 10.64898/2026.08.07.743504 medRxiv
Top 0.2%
5.3%
Show abstract

Single cell RNA-sequencing (scRNA-Seq) data are typically represented as cell-by-gene count matrices, which capture the expression of each gene as detected in the sampled cells; often a heterogeneous population of multiple different cell types or cell states. Almost all scRNA-Seq analysis workflows have a gene selection step prior to applying clustering algorithms which helps remove genes with low variability and hence reduce the high-dimensional gene space. A de-facto method for achieving selection of highly variable genes (HVG) uses dispersion and mean expression scores to evaluate the variability of each individual gene. However, methods based on direct mean-to-variance relationship for gene selection often suffer from susceptibility to variance instability and arbitrary determination of the optimal number of genes to use in downstream analysis tasks, additionally, they often prioritize genes with low abundance but high variance. Here, we propose an innovative method for selecting highly variable genes that is not based on mean to variance ratios: "Principal Genes (PG)" method; it utilizes the rotations (or loadings) from Principal Component Analysis (PCA) to calculate a novel variability score per gene that we name "Gene Principal Score (GPS)". GPS helps evaluate the genes based on their contribution in the PCA rotations and hence ranks the genes according to their variability from highest to lowest variable genes. For efficient implementation we utilize Augmented Implicitly Restarted Lanczos Bidiagonalization methods to efficiently obtain Principal Components (PCs) associated with the largest variance. Genes with the highest GPS score, i.e. Principal Genes, can then be used for downstream analysis tasks, especially the clustering step. To test the performance of our highly variable gene identification method, we use several validation strategies, including clustering of labeled single cell RNA-Seq data (i.e. data with known ground truth cell type labels). Furthermore, we measure the performance of our method against dispersion-based highly variable gene (HVG) selection approaches. We use several validation metrics, including sensitivity and adjusted rand index scores for clustering based on genes selected using our method against genes selected using HVG; and our validation datasets include six real labeled single cell RNA-Seq datasets. Our findings show that our new method, Principal Genes, is comparable and often favorable in performance in selecting highly variable genes and achieves ultra-fast gene selection from PCA results.

8
Sequence-Derived Representations versus Pfam-Domain Content for Biosynthetic Gene Cluster Retrieval

Urokov, R.; Khan, A.; Eshboyev, F.; Asadov, D.; Rahman, S.; Kushokova, D.

2026-08-22 bioinformatics 10.64898/2026.08.21.746127 medRxiv
Top 0.2%
5.2%
Show abstract

Retrieving BGCs related to those of a known producer can be regarded as a representation-learning objective. We hypothesize that ESM-2 sequence-derived representations of BGCs can improve retrieval beyond the Pfam-domain content metric. Our toolkit is the following: group-disjoint train, validation, and test assignments, validation-frozen model selection, [fi]ve seeds, and family-level paired inference. Of 6,953 atlas BGCs from 182 deduplicated Streptomyces griseus genome accessions, 5,325 silver-labeled BGCs are split into 98 training, 21 validation, and 21 test reference groups. Of the test reference groups, 16 are eligible for retrieval diagnostics. Pfam Jaccard scored Recall@50 of 0.8788, while Pfam-augmented BGC-SetNet scored 0.8472. The combination of ESM and Pfam-augmented BGC-SetNet scored 0.8769. A weighted Pfam Jaccard obtained a slightly higher score of 0.8789, which has a negligible difference compared to unweighted Pfam accard. Our results do not support the claim that sequence-derived representations can recover alternative biosynthetic pathways on this benchmark. Instead, explicit Pfam remains the major signal for this objective. Our results de[fi]ne the curation and pathway-level validation processes that are necessary for a more robust biological test.

9
Degree-ranked gene lists omit the cross-module connectors, and a partition-free centrality recovers them

Qun, Z.; Huaizheng, Z.; Yuxin, Z.; Jieying, B.; Tan, S.

2026-08-21 bioinformatics 10.64898/2026.08.10.743862 medRxiv
Top 0.2%
4.0%
Show abstract

Network centrality is the workhorse of gene prioritisation, yet what a ranking omits is rarely audited. Scoring each selection against an annotation-count-matched maximum-entropy reference--asking whether a selected gene set covers the genomes functional space or collapses it-reveals that the criterion in standard use has a measurable blind spot in exactly the class it is meant to surface. Degree, the most widely used criterion, returns the cross-module bridges that are also locally dominant--connector hubs--and omits the non-hub connectors: where 26% of the genome occupies these coordinating roles, a degree-ranked list holds 18% and an EDVS-ranked list 55%, and degrees top-1% collapses functional coverage below the reference on all five networks tested. We repurpose EDVS (Entropy of Degree-Vector Sums), an information-theoretic diversity measure, as an annotation-free, partition-free centrality that recovers this omitted class. The coverage it preserves is carried by cross-module participation P, which cannot be computed without a community partition; EDVS matches P-level coverage on all five networks using none, and retains 0.84 of its selection under edge perturbation that leaves partition-based selections at 0.21-0.46. The deficit is general: the collapse holds in the same direction on the two networks built without functional annotation (0.5-1.1 bit; co-expression, physical interaction) as on the three supervised by it (1.6-3.3 bit; RiceNet, AraNet, STRING), so supervision amplifies it rather than creates it. The remedy is bounded: EDVS ceases to preserve coverage on the sparse physical-interaction network. And the class EDVS isolates is organizational, not an importance signal: pre-registered probes--essentiality, transcription-factor identity, tissue-specificity, date/party-hub character, phenotype co-localisation--return null or reversed throughout. The conclusive ones are equivalent to their degree-matched nulls within {+/-}5 percentage points (demonstrated, not merely undetected), and the classical coupling of centrality to importance itself holds only network-dependently. Author SummaryGenes rarely act alone: many diseases and agricultural traits are shaped by genes that coordinate several biological processes rather than specialising in one. The standard way to find such genes in a network of gene interactions is to count each genes connections--its "centrality"--and rank genes by that count. We show this standard approach has a blind spot: it favours genes that dominate one process over genes that quietly bridge several processes without dominating any, and this blind spot appears across rice, thale cress, and yeast gene networks. We repurpose a diversity measure from an unrelated field (originally used to compare citation patterns) as a new way to rank genes that finds these bridging genes from network structure alone, without needing gene-function annotations--which are themselves incomplete and biased toward well-studied genes--or a prior, unstable step of splitting the network into modules. We are careful to show where the new approach also falls short: on sparse, noisy networks it stops working, and the genes it recovers are not shown to be more biologically important than other genes, only differently positioned. What that position is for is a question this work leaves open.

10
Design-informed Size Factor Estimation

Pocuca, T.; Pare, G.; Bolker, B. M.

2026-08-22 bioinformatics 10.64898/2026.08.13.744630 medRxiv
Top 0.2%
3.4%
Show abstract

Accurate normalization is essential for differential expression analysis of RNA-sequencing data. Popular normalization methods such as the median-of-ratios and trimmed mean of M-values do not leverage information from the experimental design. This may be inefficient in experiments with large-scale systematic expression changes or complex designs. Here, we introduce design-informed size factor estimation (disize), a normalization method that uses information from the experimental design to improve accuracy. disize uses a modified generalized linear mixed model to robustly distinguish between biological signal and sample-specific size factors. We also propose a mechanistically justified data-generating process for RNA-sequencing counts that is derived from previous models of transcription and sequencing. Through simulations based on this data-generating process and validating on true RNA-seq data, we show that disize recovers size factors more accurately than existing methods, particularly in challenging scenarios with low gene expression and a high proportion of differentially expressed genes; this in turn improves downstream analysis. disize provides a robust and accurate approach to normalization, highlighting the significant benefits of integrating experimental design information directly into normalization for transcriptomic datasets. Author summaryIn transcriptomic analysis, normalization adjusts for technical biases arising from library preparation and sequencing. Methods implemented in widely used packages like DESeq2 and edgeR ignore information in the experimental design during normalization. Incorporating information from the experimental design into a normalization method has the potential to yield more accurate results. To do this, we developed a new method, design-informed size factor estimation (disize), that uses a statistical model to jointly account for the biological signal defined by the design and the sample-specific batch effect. By separating the biological variation into its components, disize can more robustly estimate the batch effect. To validate our approach, we constructed a flexible simulation framework relying on a mechanistically justified data-generating process for RNA-seq data. Our benchmarks on both simulated and true RNA-seq data show that disize recovers the true size factors more accurately than existing methods, particularly in challenging scenarios with low counts or a high proportion of differentially expressed genes. This improved normalization yields more reliable downstream results in differential expression analysis.

11
Benchmarking single-cell foundation models in a zero-shot setting

Gaballa, Y.; Ahmed, S.; Abdelaal, T.

2026-08-07 bioinformatics 10.64898/2026.08.03.739553 medRxiv
Top 0.3%
3.2%
Show abstract

Single-cell foundation models have recently emerged as a promising approach for learning general- purpose representations from large-scale transcriptomic data. These models are trained on millions of cells and are designed to transfer their learned representations to a wide range of downstream tasks. However, their practical benefits compared to traditional approaches are still not fully understood. This study evaluates four foundation models, namely scGPT, SCimilarity, UCE, and Transcriptformer, across four downstream tasks: cell type annotation, human data integration, cross-species data integration, and protein expression prediction. Embeddings generated by each model were assessed using multiple public single-cell datasets and compared against conventional machine learning baselines. Performance was measured using task-specific evaluation metrics, including classification, integration, and regression metrics. The results showed that foundation model embeddings did not consistently outperform traditional approaches. In the cell type annotation task, baseline methods achieved the strongest performance across most datasets. For protein expression prediction, however, embeddings from the foundation models generally produced more accurate predictions than the baseline, with SCimilarity achieving the lowest prediction error and Transcriptformer obtaining the highest correlation scores. In the data integration task, all foundation models produced moderate results, while scVI (the baseline) achieved the strongest integration performance. Overall, the results suggest that current single-cell foundation models provide useful representations for some downstream tasks in zero-shot conditions but do not yet offer a universal replacement for task-specific methods. Their effectiveness remains dependent on the application and evaluation setting.

12
Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles

Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.

2026-08-25 bioinformatics 10.64898/2026.08.21.745839 medRxiv
Top 0.4%
2.7%
Show abstract

Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.

13
Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach

Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.

2026-09-01 bioinformatics 10.64898/2026.08.27.747483 medRxiv
Top 0.4%
2.5%
Show abstract

Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.

14
ECHO: A lightweight tool for inferring missing case counts from pathogen phylogenies

Doig, R.; Colijn, C.

2026-08-11 epidemiology 10.64898/2026.08.09.26360045 medRxiv
Top 0.4%
2.4%
Show abstract

Timed phylogenetic trees express the evolutionary history of a pathogen outbreak in units of time, providing an estimate of the elapsed time across the shared ancestry of a set of taxa. By combining this elapsed time with known information about the epidemiology of a disease, we can relate the total branch length to the total number of cases related to the phylogeny. This gives information about the number of unsequenced cases that are related to the phylogeny. We call these ``cryptic'' cases. We present ECHO (Estimation of Cryptic Hosts from Outbreak trees), a collection of three lightweight estimators of the number of cryptic cases in a phylogeny. ECHO is agnostic to the form of the sampling process, making it robust to a variety of forms of sampling heterogeneity. We demonstrate ECHO's baseline accuracy and its robustness to heterogenous sampling frameworks through simulation. Additionally, we apply ECHO to measles virus sequences that were collected during an outbreak in the USA in 2021. ECHO is able to recover the number of cryptic cases with a reasonable degree of accuracy both in simulation and in practice. We discuss the contexts in which ECHO is most applicable, and the interpretation of its estimates.

15
Dynamic Hierarchical Interleaved Bloom Filter: An Updatable Index for Large-Scale Fast Sequence Search

Seiler, E.; Willemsen, M.; Piro, V. C.; Reinert, K.

2026-08-30 bioinformatics 10.64898/2026.08.26.747224 medRxiv
Top 0.4%
2.3%
Show abstract

Motivation: A continued decrease in sequencing costs has facilitated the exponential increase in available sequencing data, with public databases like the European Nucleotide Archive (ENA) and Sequence Read Archive (SRA) reaching well in the order of petabases. This has been the incentive to develop more scalable tools for common bioinformatics tasks. One such task is the approximate searching of short sequence patterns like genes or reads in reference data sets. In recent years, a variety of indexing data structures have been proposed for searching large sequencing databases. The state-of-the-art index, the Hierarchical Interleaved Bloom Filter (HIBF) was first-in-class to index one million samples. To be useful for expanding repositories, it must be extended to support dynamic updates. Results: In this paper, we introduce a scalable and updatable sequence-search index by extending the HIBF with partial rebuilding to support efficient updates. We demonstrate the Dynamic HIBF's capacity for large-scale data by iteratively creating an index from over 100 TB of compressed reads across more than 39,000 full human RNA-Seq samples, updated in consecutive batches of 100. To benchmark against state-of-the-art tools, we evaluated incremental performance on a subset of 5,000 samples sub-sampled to 1% of their original read depth. In this comparative setting, the dynamic HIBF completed the sequential insertion of all 5,000 samples within 5 hours--24 to 65 times faster than competing methods and twice as fast as the static HIBF.

16
Learning Shared Residue Backgrounds and Modification-Specific Offsets for PTM Site Prediction

Pokharel, S.; Bhusal, B.

2026-08-11 bioinformatics 10.64898/2026.08.05.743079 medRxiv
Top 0.4%
2.1%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWPost-translational modifications (PTMs) are chemical changes added to proteins after translation. These changes affect protein function and regulation, and their disruption is linked to disease-associated mechanisms. Because experimentally validating all possible modification sites is impractical, many computational predictors have been developed for PTM site prediction. In this work, we study whether a shared model can represent common residue-background patterns while learning modification-specific background-to-positive offsets. This framing is especially relevant for residues such as lysine (K), which can be acetylated, ubiquitinated, methylated, or sumoylated depending on the surrounding protein context. We propose an anchor-guided rectified flow matching framework for multi-type PTM site prediction from protein language model embeddings. For each PTM-residue pair, the model builds residue-background anchors from PTM-compatible unannotated residues and positive anchors from experimentally annotated modified residues. Given a candidate residue and target modification type, the model compares the residue embedding with these anchor sets and uses a rectified flow module to estimate a modification-conditioned background-to-positive offset. This offset is combined with anchor-based features and used for site scoring. We evaluate the framework on a dbPTM-derived benchmark covering six commonly studied PTMs: phosphorylation, acetylation, ubiquitination, methylation, sumoylation, and N-linked glycosylation. In the shared-model setting, our approach achieves a macro AUPRC of 0.4195, improving over the gated multi-anchor baseline of 0.4154, while independently trained per-modification models achieve 0.4353. These results suggest that multi-type PTM prediction can be modeled within a single shared framework by combining residue-background anchors with modification-conditioned offset features.

17
LRSPAT: A low-rank framework for spatial omics statistics

Frost, H. R.

2026-08-31 bioinformatics 10.64898/2026.08.26.747291 medRxiv
Top 0.5%
2.1%
Show abstract

We describe LRSPAT (low-rank spatial toolkit), a fast and memory-efficient framework for approximating measures of spatial association for high-dimensional data. While LRSPAT can be applied to any multivariate spatial dataset, development was motivated by the computational challenge of identifying spatially variable genes in high-resolution spatial transcriptomics (ST) data generated by technologies such as 10x Visium HD, Xenium and Atera. LRSPAT leverages a truncated SVD of the expression data and a thresholded spatial weights matrix to perform reduced-rank reconstruction of spatial statistics in the quadratic form family, including global and local versions of Moran's I, Geary's C, and Getis-Ord G. A regularization approach is leveraged to account for the inflated null distribution of spatial statistics computed on latent variables. By performing key operations on the low-dimensional embeddings, LRSPAT is orders of magnitude faster than standard implementations with significantly lower memory requirements. Because the low-rank approach denoises and desparsifies ST data, LRSPAT is also more accurate than standard techniques at identifying genes with true spatial expression patterns. The dramatic improvements in execution time and memory consumption enable the genome-wide analysis of spatially variable genes (SVGs) and exploration of the full range of hyperparameters including spatial scale, distance metric, and embedding rank. This preprint outlines the background and mathematical details of the approach with limited preliminary results and a short conclusion.

18
A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences

Motta, J. A.; Motta, M. d. M.; Fernandez, C.

2026-08-20 bioinformatics 10.64898/2026.08.16.745093 medRxiv
Top 0.5%
1.9%
Show abstract

In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.

19
How do noise and precision contraints jointly determine the minimum sketch size for fixed-size MinHash Jaccard estimation?

Ebou, A. E. T.; KOUA, D. K.

2026-08-20 bioinformatics 10.64898/2026.08.15.744971 medRxiv
Top 0.5%
1.8%
Show abstract

MinHash-based genome comparison is governed by two statistical constraints: a noise floor on k-mer size, controlling chance collisions between unrelated sequences, and a precision floor on sketch size, controlling uncertainty in the estimated Jaccard similarity. These constraints are best characterized for fixed-size bottom-sketch MinHash, the classical construction underlying tools such as Mash and still the standard basis for comparing genomes of similar size. Nevertheless, while, the noise floor has an established closed-form solution, sketch size is commonly selected using fixed defaults or heuristic choices independent of genome length and target precision. Moreover, recent work on scaled (FracMinHash) sketching has highlighted the difficulty of obtaining an analogous closed-form confidence interval for the Jaccard similarity. Here, we show that under the fixed-size bottom-sketch MinHash model, the shared-hash count follows a binomial model, which lets classical proportion-estimation theory be applied directly. Combining this with Fofanov's k-selection criterion and the exact identity-Jaccard relationship yields a single closed-form design rule, s* (L, {rho}* , ), that unifies the noise and precision floors into one operating envelope. The resulting design rule was validated against exact Clopper-Pearson interval inversion, idealized binomial sampling, realistic overlapping k-mer simulation and 54 bacterial genome pairs spanning species-to-order taxonomic levels. In realistic simulations, empirical coverage remained within 1-3 percentage points of the idealized reference in 14 of 15 tested cells. In real genomes, the classical identity-Jaccard relationship showed increasing positive bias with taxonomic divergence, from a median of -1.4% within species to +27% at genus and +77% at family level. We further show that s* (L) is not smooth in genome length but a discontinuous, previously unreported staircase caused by the ceiling function used for k-mer selection, with practical consequences concentrated at specific genome-size boundaries.

20
Ancestral Sequences Cannot be Accurately Reconstructed via Interpolation in a Variational Autoencoder's Latent Space

Gorstein, E.; Tang, M.; Bruzzone, H.; Solis-Lemus, C.

2026-09-01 evolutionary biology 10.1101/2025.11.19.689264 medRxiv
Top 0.6%
1.5%
Show abstract

Standard methods for ancestral sequence reconstruction (ASR) rely on substitution models for the residues in a biological sequence and assume independent evolution across these sites, ignoring the epistatic interactions that shape molecular evolution. In contrast, deep learning models like variational autoencoders (VAEs) can learn low-dimensional representations ("embeddings") of sequences in a protein family that may implicitly handle these dependencies, raising the possibility of performing more accurate ASR by interpolating between extant sequence embeddings within the VAE's latent space. In this study, we test this hypothesis by developing and evaluating a VAE-based ASR pipeline. Benchmarking this approach against established likelihood-based and parsimony methods using various simulations of protein evolution, including scenarios with and without epistasis, we find that the VAE-based approach is consistently and significantly outperformed by standard methods, even in epistatic regimes where it was hypothesized to have an advantage. We further show that this failure is not due to a lack of phylogenetic structure in the latent space, which does contain evolutionary signal. Rather, the primary limitation is the information loss inherent to the autoencoding process: the VAE's decoder cannot generate sequences with sufficient fidelity for the precise demands of ASR.